Papers by Yassine El Kheir

4 papers
MorphBPE: Morphology-Aware Tokenization for Efficient LLM Training (2026.findings-acl)

Copied to clipboard

Challenge: Tokenization is a key design choice in modern NLP systems and a critical bottleneck for multilingual Large Language Models.
Approach: They propose a tokenization extension that constrains merge operations to respect morpheme boundaries while preserving inference.
Outcome: The proposed tokenization improves morphological coherence and language model cross-entropy in four languages.
LAraBench: Benchmarking Arabic AI with Large Language Models (2024.eacl-long)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have significantly influenced the landscape of language and speech research.
Approach: They used GPT-3.5-turbo, GPT-4, BLOOMZ, Jais-13b-chat, Whisper, and USM to tackle 33 distinct tasks across 61 datasets.
Outcome: The proposed model outperforms SOTA models in zero-shot learning, with a few exceptions.
Beyond Orthography: Automatic Recovery of Short Vowels and Dialectal Sounds in Arabic (2024.acl-long)

Copied to clipboard

Challenge: Existing algorithms for recognizing borrowed and dialectal sounds are limited to Arabic, a dialect-rich language containing more than 22 major dialects.
Approach: They propose a framework to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages that extends beyond its standard orthographic sound sets.
Outcome: The proposed framework improves character error rate by 7% with only one and half hours of training data compared to the baseline.
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection (2025.findings-naacl)

Copied to clipboard

Challenge: Existing algorithms for audio deepfake detection are based on layer-wise analysis of self-supervised learning (SSL) models.
Approach: They conduct a layer-wise analysis of self-supervised learning (SSL) models for audio deepfake detection across diverse contexts.
Outcome: The proposed models achieve competitive equal error rate (EER) scores even when employing a reduced number of layers.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations